Papers with Language documentation
Joint Word and Morpheme Segmentation with Bayesian Non-Parametric Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | Language documentation often requires segmenting transcriptions of utterances into words and morphemes . a long tradition of nonparametric Bayesian models is used to handle these tasks . |
| Approach: | They propose a Bayesian model for simultaneously segmenting utterances at two levels . they use two under-resourced languages to better understand the value of weak supervision . |
| Outcome: | The proposed model can be used to identify language documents with weak supervision. |
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing resources for standardized, easily accessible IGT data limit their applicability to linguistic research. |
| Approach: | They compile the largest existing corpus of interlinear glossed text data from a variety of sources and use it to generate annotated text. |
| Outcome: | The proposed model outperforms SOTA models on monolingual corpora by 6.6%. |